Papers with Language model evaluation
BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for evaluating language models are brittle, corpus-level perplexities are vague, and the choice of benchmarks is endless. |
| Approach: | They propose a method that uses contextual embeddings to find fine-grained features of text where one model outperforms another. |
| Outcome: | The proposed method extracts features that demonstrate differences with respect to ease of generation between two language models. |